Near Duplicate Detection fields

After the job completes following the promotion job, the system populates the following Near Duplicate Detection related fields for non-email documents.

Field name Description
Near Duplicate ID

Displays a numeric value. Documents with the same are nearly identical to each other and originate from the same reference copy.

Pivot

Indicates the reference document used to identify identical documents. When multiple documents share the same , the first document processed becomes the pivot. Documents that are not identical to any other documents are also marked as Pivot.

Similarity

Displays the percentage score that indicates how similar a document is to the pivot document. For a pivot document, this field displays 100.

When you run the NDT document Action, the system populates the above fields for the selected documents. The system prefixes the Pivot and Similarity fields with the NDT job name.

The following table provides the system behavior on populating the fields when you run Near Duplicate Detection.

Scenario

System Behavior

First time NDT run

Populates Pivot, Near Duplicate ID, and Similarity fields.

Run NDT again on a subset without original pivot

Identifies a new pivot document for that subset.

Documents already in a group

Reuses the existing Near Duplicate ID.

NDT already run during promotion and you run the NDT document action

Only Similarity and Pivot fields are recalculated for the selected set.